Papers with visual dialog models

2 papers
CLEVR-Dialog: A Diagnostic Dataset for Multi-Round Reasoning in Visual Dialog (N19-1)

Copied to clipboard

Challenge: Visual Dialog is a multimodal task of answering a sequence of questions grounded in an image.
Approach: They construct a dialog grammar that is grounded in the scene graphs of the images from the CLEVR dataset and use it to benchmark performance of standard visual dialog models.
Outcome: The proposed model is based on a large diagnostic dataset for studying multi-round reasoning in visual dialog.
Learning to Ground Visual Objects for Visual Dialog (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to ground visual objects are inadequate for visual dialog . a posterior distribution is inferred from context and questions, while posterior distributions are used to facilitate visual objects grounding.
Approach: They propose a method to learn to ground visual objects for visual dialog using prior and posterior distributions over visual objects to facilitate visual objects grounding.
Outcome: The proposed approach improves the existing models in generative and discriminative settings by a significant margin.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations